Back

NAR Genomics and Bioinformatics

Oxford University Press (OUP)

Preprints posted in the last 30 days, ranked by how well they match NAR Genomics and Bioinformatics's content profile, based on 242 papers previously published here. The average preprint has a 0.17% match score for this journal, so anything above that is already an above-average fit.

1
Method Choice, Not Biology, Determines In Silico Perturbation Results: A Systematic Evaluation of Eight Methods Across Four Datasets

Wenjie, G.; Wu, S.; Hu, G.; Yang, Z.; Wang, Z.; Cai, J.; Mao, J.

2026-08-19 bioinformatics 10.64898/2026.08.11.744106 medRxiv
Top 0.1%
21.5%
Show abstract

Most in silico perturbation methods for single-cell transcriptomics have been validated only on individual datasets, leaving their reliability and generalizability unknown. Through systematic cross-method, cross-dataset benchmarking of eight methods spanning six mathematical frameworks across four datasets, we find that six of eight methods--including widely used VAE-based and tensor decomposition approaches--fail to produce detectable transcription factor (TF)-to-pathway signals. Only CellOracle and DDIM consistently detected TF-to-glycolysis directional regulation. Cross-pathway analysis in PBMC monocytes revealed biologically coherent TF-pathway associations beyond glycolysis (SPI1[->]glycolysis 4.4x enrichment, FOS[->]AP-1 targets 4.4x), with SOX9 serving as a biological specificity control (no pathway enrichment). Method choice alone could reverse biological conclusions: DDIM and scTenifoldKnk rankings were significantly anti-correlated ({rho}=-0.811, p=0.027). CRISPRi Perturb-seq validation in K562 cells confirmed TF knockdown suppresses glycolysis gene expression (JUN {delta}=-1.72, CEBPB {delta}=-1.59, SPI1 {delta}=-1.57, FOS {delta}=-0.70), but CellOracle-predicted perturbation directions did not match experimental directions (40.9% agreement, not different from chance), revealing a fundamental gap between steady-state correlation and causal perturbation. Diagnostic analyses using VAE latent space profiling, correlation distribution comparison, and gene-gene graph analysis identified distinct failure modes in unsuccessful methods: VAE latent space competition (STAT3 signal-to-noise 0.44 vs. SPI1 4.25), correlation noise (TF-glycolysis |r|=0.038 indistinguishable from background |r|=0.047), and graph non-specificity (0.84x enrichment). A controlled ablation experiment showed that adding a GRN prior to DDIM did not improve target recall (delta=0 for all TFs), confirming that performance differences are multi-factorial. These findings establish preliminary guidance for method selection, including cross-pathway validation, direction-aware benchmarking, and minimum data requirements ([≥]500 cells, [≥]1,000 HVGs).

2
How do noise and precision contraints jointly determine the minimum sketch size for fixed-size MinHash Jaccard estimation?

Ebou, A. E. T.; KOUA, D. K.

2026-08-20 bioinformatics 10.64898/2026.08.15.744971 medRxiv
Top 0.1%
14.5%
Show abstract

MinHash-based genome comparison is governed by two statistical constraints: a noise floor on k-mer size, controlling chance collisions between unrelated sequences, and a precision floor on sketch size, controlling uncertainty in the estimated Jaccard similarity. These constraints are best characterized for fixed-size bottom-sketch MinHash, the classical construction underlying tools such as Mash and still the standard basis for comparing genomes of similar size. Nevertheless, while, the noise floor has an established closed-form solution, sketch size is commonly selected using fixed defaults or heuristic choices independent of genome length and target precision. Moreover, recent work on scaled (FracMinHash) sketching has highlighted the difficulty of obtaining an analogous closed-form confidence interval for the Jaccard similarity. Here, we show that under the fixed-size bottom-sketch MinHash model, the shared-hash count follows a binomial model, which lets classical proportion-estimation theory be applied directly. Combining this with Fofanov's k-selection criterion and the exact identity-Jaccard relationship yields a single closed-form design rule, s* (L, {rho}* , ), that unifies the noise and precision floors into one operating envelope. The resulting design rule was validated against exact Clopper-Pearson interval inversion, idealized binomial sampling, realistic overlapping k-mer simulation and 54 bacterial genome pairs spanning species-to-order taxonomic levels. In realistic simulations, empirical coverage remained within 1-3 percentage points of the idealized reference in 14 of 15 tested cells. In real genomes, the classical identity-Jaccard relationship showed increasing positive bias with taxonomic divergence, from a median of -1.4% within species to +27% at genus and +77% at family level. We further show that s* (L) is not smooth in genome length but a discontinuous, previously unreported staircase caused by the ceiling function used for k-mer selection, with practical consequences concentrated at specific genome-size boundaries.

3
pysigscore: gene signatures scoring across bulk and single-cell transcriptomics

Giacomello, T.; Mazzara, S.; Abbruzzese, G.; Barberis, A.; tangherloni, a.; Buffa, F. M.

2026-08-09 bioinformatics 10.64898/2026.08.04.742537 medRxiv
Top 0.1%
12.5%
Show abstract

SummaryHigh-throughput transcriptomics has made gene signatures central to interpreting gene expression data, with applications in diagnosis, prognosis, and prediction. Quantifying signature activity and assessing its robustness remain challenging because scoring methods primarily rely on various assumptions, and no single approach is universally optimal. Here, we present pysigscore, a Python framework for gene set scoring in bulk and single-cell RNA-seq data. pysigscore integrates 18 built-in scoring methods with a fully customisable scorer, allowing users to define and benchmark new scoring functions. It also provides reliability analyses, including p-value estimation and leave-one-out experiments, to assess the significance of scores and gene-level contributions. We validated pysigscore on the CCLE, TCGA, and PBMC datasets, recovering the expected enrichment in liver, hypoxia, inflammatory, and cell-cycle signatures. Availability and ImplementationSource code is available at https://github.com/bioinformatics-hub/pysigscore. Contact: tommaso.giacomello@phd.unibocconi.it, francesca.buffa@unibocconi.it Supplementary informationSupplementary data are available at Bioinformatics online.

4
Novel biologically relevant small RNA-sequencing alignment tool LevenMap for alignment to database of non-coding RNAs

Dlugas, H.; Dyson, G.; Dombkowski, A.; Kim, Y.; Gurdziel, K.; Boerner, J. L.; Bock, C.

2026-08-21 bioinformatics 10.64898/2026.08.14.742100 medRxiv
Top 0.1%
12.4%
Show abstract

A crucial aspect of the bioinformatics workflow in small RNA-sequencing is the alignment of reads to a database of reference ncRNAs. Alignment algorithms such as Bowtie, Burrows-Wheeler Aligner (BWA), and Spliced Transcripts Alignment to a Reference (STAR) - which are designed for aligning reads to a reference genome - are typically used. Aligning short RNA-sequenced reads to a database of non-coding RNAs (ncRNAs) is fundamentally a different task than aligning longer reads to a genome due to ncRNAs (i) having roughly the same number of nucleotides as the reads being aligned and (ii) being subsequences of other ncRNAs. To account for these differences, we developed the novel alignment algorithm LevenMap. Of all reads which exactly matched a reference ncRNA in a publicly available dataset, LevenMap aligned 100.0% of them to their respective ncRNA while all other aligners mapped less than 40% of these reads to their corresponding ncRNA. Furthermore, the mean ratio (length of read) / (length of corresponding reference ncRNA) of all aligned reads was 1.0 and 0.998 for LevenMap with at most zero and one mismatch(es) allowed, respectively; this ratio was no more than 0.51 for all other aligners. Overall, LevenMap is designed to account for the nuances of aligning small RNA-sequencing data to a database of reference ncRNAs and yields more biologically relevant counts compared to traditional aligners in this context. LevenMap is free and publicly available on GitHub: https://github.com/hdlugas/LevenMap.

5
CoTRA: an integrated R/Shiny framework for transparent bulk and single-cell RNA-seq analysis

Seemab, U.; Vainionpaa, K.; Tanoli, Z.; Leinonen, H. O.

2026-08-28 bioinformatics 10.64898/2026.08.25.747017 medRxiv
Top 0.1%
12.0%
Show abstract

Bulk and single-cell RNA sequencing (scRNA-seq) have become essential for investigating disease mechanisms and identifying diagnostic biomarkers. However, the growing volume of transcriptomic data remains difficult to reuse efficiently for many researchers. Downstream analysis often requires multiple statistical, visualization, and reporting tools, creating fragmented workflows that reduce transparency and reproducibility, particularly when analyzing scRNA-seq data. To address this gap, we developed CoTRA (Comprehensive Toolbox for RNA-seq Analysis), an open-source R/Shiny package for bulk and scRNA analysis. CoTRA integrates established methods into modular workflows, exposes parameters, and offers alternatives at selected stages. It supports bulk RNA-seq quality assessment, differential expression, annotation, enrichment, and reporting, as well as scRNA quality control, dimensionality reduction, clustering, marker identification, cell-type annotation, differential abundance, trajectory inference, pathway activity, and cell-cell communication. CoTRA runs on workstations or HPC environments without mandatory external data submission and was tested on Linux, Windows, and macOS. Compared with 14 other platforms for bulk RNA-seq/scRNA-seq, CoTRA supported 46 of 49 predefined functionality criteria. Tool validation using published rd10 retinal bulk RNA-seq identified 1,947 shared differentially expressed genes with concordant direction and strong log2 fold-change agreement. A retinal scRNA-seq case study demonstrated appropriate clustering, cell-type resolved analysis, and pathway activity scoring. CoTRA provides a graphical environment for bulk and single-cell RNA-seq analysis while retaining parameter transparency, methodological flexibility, and reproducible outputs. Strong concordance with the published bulk RNA-seq analysis supports the workflow consistency, while the single-cell case study demonstrates its applicability to advanced scRNA-seq analysis. The source code is freely available at https://github.com/UmairSeemab/CoTRA.

6
IsoAtlas: Visual interpretation of known and novel transcript isoforms using population-scale long-read evidence

Zheng, X.; Sedlazeck, F. J.

2026-08-21 bioinformatics 10.64898/2026.08.12.744438 medRxiv
Top 0.1%
12.0%
Show abstract

Long-read RNA sequencing has revealed extensive transcript diversity, but newly observed isoforms remain difficult to interpret beyond their classification as known or novel. Here, we present IsoAtlas, an interactive multispecies database for visual exploration and population-scale interpretation of transcript isoforms using 1, 035 human and 414 mouse uniformly processed long-read RNA-sequencing samples. Users can search annotated genes and transcripts or submit novel transcript models in GTF format, visualize their structures, and examine sample-level support, prevalence, expression, tissue and disease context, and sequencing-platform evidence. IsoAtlas integrates structurally equivalent transcripts across GENCODE, RefSeq and CHESS, consolidating evidence that would otherwise be distributed across annotation-specific identifiers. It further links corresponding human and mouse transcript models, enabling users to assess cross-species conservation and enabling users to assess cross-species conservation and inform the suitability of mouse models for isoform-specific studies. IsoAtlas can also evaluate arbitrary user-supplied transcript structures directly against accumulated long-read evidence. IsoAtlas therefore complements established reference annotations with an extensible evidence layer that connects transcript structure to population prevalence, biological context and cross-species support. IsoAtlas is freely available at https://www.isoatlas.org/. Graphic abstract O_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=83 SRC="FIGDIR/small/744438v1_ufig1.gif" ALT="Figure 1"> View larger version (21K): org.highwire.dtl.DTLVardef@89bd72org.highwire.dtl.DTLVardef@f48bbforg.highwire.dtl.DTLVardef@102ae37org.highwire.dtl.DTLVardef@fbbc2b_HPS_FORMAT_FIGEXP M_FIG C_FIG

7
Sequence-Derived Representations versus Pfam-Domain Content for Biosynthetic Gene Cluster Retrieval

Urokov, R.; Khan, A.; Eshboyev, F.; Asadov, D.; Rahman, S.; Kushokova, D.

2026-08-22 bioinformatics 10.64898/2026.08.21.746127 medRxiv
Top 0.2%
9.8%
Show abstract

Retrieving BGCs related to those of a known producer can be regarded as a representation-learning objective. We hypothesize that ESM-2 sequence-derived representations of BGCs can improve retrieval beyond the Pfam-domain content metric. Our toolkit is the following: group-disjoint train, validation, and test assignments, validation-frozen model selection, [fi]ve seeds, and family-level paired inference. Of 6,953 atlas BGCs from 182 deduplicated Streptomyces griseus genome accessions, 5,325 silver-labeled BGCs are split into 98 training, 21 validation, and 21 test reference groups. Of the test reference groups, 16 are eligible for retrieval diagnostics. Pfam Jaccard scored Recall@50 of 0.8788, while Pfam-augmented BGC-SetNet scored 0.8472. The combination of ESM and Pfam-augmented BGC-SetNet scored 0.8769. A weighted Pfam Jaccard obtained a slightly higher score of 0.8789, which has a negligible difference compared to unweighted Pfam accard. Our results do not support the claim that sequence-derived representations can recover alternative biosynthetic pathways on this benchmark. Instead, explicit Pfam remains the major signal for this objective. Our results de[fi]ne the curation and pathway-level validation processes that are necessary for a more robust biological test.

8
Degree-ranked gene lists omit the cross-module connectors, and a partition-free centrality recovers them

Qun, Z.; Huaizheng, Z.; Yuxin, Z.; Jieying, B.; Tan, S.

2026-08-21 bioinformatics 10.64898/2026.08.10.743862 medRxiv
Top 0.2%
9.8%
Show abstract

Network centrality is the workhorse of gene prioritisation, yet what a ranking omits is rarely audited. Scoring each selection against an annotation-count-matched maximum-entropy reference--asking whether a selected gene set covers the genomes functional space or collapses it-reveals that the criterion in standard use has a measurable blind spot in exactly the class it is meant to surface. Degree, the most widely used criterion, returns the cross-module bridges that are also locally dominant--connector hubs--and omits the non-hub connectors: where 26% of the genome occupies these coordinating roles, a degree-ranked list holds 18% and an EDVS-ranked list 55%, and degrees top-1% collapses functional coverage below the reference on all five networks tested. We repurpose EDVS (Entropy of Degree-Vector Sums), an information-theoretic diversity measure, as an annotation-free, partition-free centrality that recovers this omitted class. The coverage it preserves is carried by cross-module participation P, which cannot be computed without a community partition; EDVS matches P-level coverage on all five networks using none, and retains 0.84 of its selection under edge perturbation that leaves partition-based selections at 0.21-0.46. The deficit is general: the collapse holds in the same direction on the two networks built without functional annotation (0.5-1.1 bit; co-expression, physical interaction) as on the three supervised by it (1.6-3.3 bit; RiceNet, AraNet, STRING), so supervision amplifies it rather than creates it. The remedy is bounded: EDVS ceases to preserve coverage on the sparse physical-interaction network. And the class EDVS isolates is organizational, not an importance signal: pre-registered probes--essentiality, transcription-factor identity, tissue-specificity, date/party-hub character, phenotype co-localisation--return null or reversed throughout. The conclusive ones are equivalent to their degree-matched nulls within {+/-}5 percentage points (demonstrated, not merely undetected), and the classical coupling of centrality to importance itself holds only network-dependently. Author SummaryGenes rarely act alone: many diseases and agricultural traits are shaped by genes that coordinate several biological processes rather than specialising in one. The standard way to find such genes in a network of gene interactions is to count each genes connections--its "centrality"--and rank genes by that count. We show this standard approach has a blind spot: it favours genes that dominate one process over genes that quietly bridge several processes without dominating any, and this blind spot appears across rice, thale cress, and yeast gene networks. We repurpose a diversity measure from an unrelated field (originally used to compare citation patterns) as a new way to rank genes that finds these bridging genes from network structure alone, without needing gene-function annotations--which are themselves incomplete and biased toward well-studied genes--or a prior, unstable step of splitting the network into modules. We are careful to show where the new approach also falls short: on sparse, noisy networks it stops working, and the genes it recovers are not shown to be more biologically important than other genes, only differently positioned. What that position is for is a question this work leaves open.

9
TxNova: recovery of recurrent unannotated intergenic splice loci from existing bulk RNA-seq alignments

Li, Z.; James, A.; Li, S.

2026-08-25 bioinformatics 10.64898/2026.08.24.746847 medRxiv
Top 0.2%
9.6%
Show abstract

Background. Reference catalogs such as GENCODE capture most stably expressed mammalian genes but may not include condition-restricted or low-abundance transcripts. Reads supporting unannotated intergenic splice junctions are generally absent from annotation-restricted gene-count matrices; transcript assemblers may reconstruct a subset as novel models, but those models are typically handled separately from the annotated gene-count universe. TxNova is a lightweight command-line tool that directly indexes these unannotated splices from existing BAM files - without an external assembler, gffcompare, or workflow manager - yielding candidate leads for bench validation rather than standalone discovery claims. Methods. TxNova retains CIGAR N junctions from STAR/HISAT2 BAM files that recur in >= 2 samples and are absent from a comprehensive annotation, clusters them into residual loci, and counts each locus alongside annotated genes in a unified matrix. Structure gates - canonical splice motif, same-strand distance, coverage valley, bridging-junction absence, minimum length - remove likely artifacts to yield structure-pass models. An optional contrast filter retains loci detected in treatment but nearly silent in control. Results. A residual splice is a recurrent, unannotated CIGAR N junction that does not overlap any annotated gene body; intronic and antisense channels are out of scope. Across four published mouse treatment arms (GSE221720, GSE166522, GSE157460, GSE193335), harvest catalogs yielded 464, 657, 789, and 594 loci. On GSE221720, 45% of loci (209/464) shared an exact intron with another series, versus 0.074% for excluded junctions; a coordinate-placement control yielded 0/5,000 matches for length-matched intergenic decoys. Masked-gene recovery reached 87.6% (176/201) overall and 99.4% (176/177) among genes with a leak junction - sensitivity rather than precision estimates. An optional contrast provides a presence/absence screen on interval TPM. Residual models are partial reconstructions requiring cloning, RACE, or targeted proteomics for confirmation. Conclusions. TxNova produces a reproducible intergenic residual-locus catalog from existing BAM files, with four downloadable mouse injury/infection catalogs. Cross-series recurrence and coordinate-based null controls support reproducibility above simple placement backgrounds but do not by themselves establish biological validity. The optional contrast step offers a practical screen for candidate leads warranting experimental validation.

10
Dynamic Hierarchical Interleaved Bloom Filter: An Updatable Index for Large-Scale Fast Sequence Search

Seiler, E.; Willemsen, M.; Piro, V. C.; Reinert, K.

2026-08-30 bioinformatics 10.64898/2026.08.26.747224 medRxiv
Top 0.2%
9.6%
Show abstract

Motivation: A continued decrease in sequencing costs has facilitated the exponential increase in available sequencing data, with public databases like the European Nucleotide Archive (ENA) and Sequence Read Archive (SRA) reaching well in the order of petabases. This has been the incentive to develop more scalable tools for common bioinformatics tasks. One such task is the approximate searching of short sequence patterns like genes or reads in reference data sets. In recent years, a variety of indexing data structures have been proposed for searching large sequencing databases. The state-of-the-art index, the Hierarchical Interleaved Bloom Filter (HIBF) was first-in-class to index one million samples. To be useful for expanding repositories, it must be extended to support dynamic updates. Results: In this paper, we introduce a scalable and updatable sequence-search index by extending the HIBF with partial rebuilding to support efficient updates. We demonstrate the Dynamic HIBF's capacity for large-scale data by iteratively creating an index from over 100 TB of compressed reads across more than 39,000 full human RNA-Seq samples, updated in consecutive batches of 100. To benchmark against state-of-the-art tools, we evaluated incremental performance on a subset of 5,000 samples sub-sampled to 1% of their original read depth. In this comparative setting, the dynamic HIBF completed the sequential insertion of all 5,000 samples within 5 hours--24 to 65 times faster than competing methods and twice as fast as the static HIBF.

11
MPGEM: A harmonized and transcriptome-complete resource for large-scale reuse of legacy human microarray data

Gupta, S.; Verma, A. K.; Jana, S.; Ahmad, S.

2026-08-25 bioinformatics 10.64898/2026.08.20.746052 medRxiv
Top 0.3%
8.0%
Show abstract

Abstract Background: Legacy microarray datasets provide an extensive record of human transcriptomic biology, but their reuse is constrained by differences in platform design, preprocessing, measurement scale, and gene coverage. Platforms measuring only subsets of genes cannot readily be integrated with higher-coverage platforms, limiting large-scale analysis and computational modeling. Results: We developed Multi-Platform Gene Expression Matrix (MPGEM), a computational framework and resource for harmonizing and completing gene-expression profiles across heterogeneous microarray platforms. MPGEM uses a Reference Quantile Distribution (RQD) and generalized Reference Subset Quantile Distribution (RSQD) framework to transform profiles with different gene coverage onto a common quantitative scale. The MPGEM Engine, a multilayer perceptron, predicts expression of unmeasured genes from genes shared across platforms. Applied to Affymetrix GPL570, GPL571, and GPL96, MPGEM uses GPL570 as a 19,320- gene reference space comprising 12,712 predictor and 6,608 target genes. The resulting resource contains 207,135 human gene-expression profiles across 19,320 genes. Evaluation using masked GPL570 profiles yielded mean sample-wise Pearson and Spearman correlations of 0.944 and 0.939, respectively, and mean gene-wise correlations of 0.830 and 0.825. The lowest-performing 5% of target genes achieved a mean Pearson correlation of 0.683. MPGEM showed comparable or higher predictive performance than baseline mean imputation and K-nearest-neighbor approaches. Conclusions: MPGEM transforms heterogeneous, partially measured legacy microarray profiles into a harmonized, transcriptome-complete representation, facilitating their reuse for large-scale transcriptomic analysis, biomarker discovery, systems biology, and machine learning. The framework, trained models, and expression resource are provided as open-source resources.

12
EnSEMBLE: a framework for enhancer-anchored pathway analysis that locks in enhancer-corroborated pathways from transcriptome sequencing data for biological validation

Zhang, L.; Gupta, A.; Wang, Y.; Sharma, R.; Lawal, B.; Hou, G.; Wang, X.-S.

2026-08-21 bioinformatics 10.64898/2026.08.17.745283 medRxiv
Top 0.3%
7.9%
Show abstract

Background Pathway discovery methods for transcriptome sequencing return tens to hundreds of redundant gene sets, and biologists often subjectively select the pathways fitting biological expectations. What is missing is not another statistical method, but a way to corroborate each candidate pathway against an independent, mechanistic line of evidence. Results We introduce EnSEMBLE (Enhancer-Set Enrichment & Mechanism-Based Linked Evidence), a tool that corroborates gene-level pathway enrichment with an orthogonal enhancer layer drawn from the same transcriptome sequencing data: active enhancers transcribe enhancer RNAs already present in standard RNA-seq, so a pathway's regulatory state can be scored from the very run that produced the gene-level signal, at no added cost. EnSEMBLE pairs pathway enrichments with Enhancer-Program Enrichment Analysis (EPEA), collapses redundant gene sets into process-level Themes, and retains only those that a concordant enhancer program corroborates. This dual-evidence requirement reduced reported signatures by >97% (hundreds of gene sets to 3-18 claims) across four datasets spanning cancer perturbations and iPSC-to-neuron differentiation. Surviving claims recovered expected biology--mesenchymal-program collapse upon SNAI1 knockout, regulatory convergence during neuronal differentiation--and named mechanisms pathway enrichments missed, including an mTOR-MYC-SPT5 elongation axis in rapamycin-treated PANC1 cells. A language AI agent performs narrative synthesis over deterministic statistics, with reproducibility enforced by temperature-zero inference and three-run consensus. We further provide enhancer over-representation analysis (eORA), mapping non-coding GWAS variants to the same programs to recover cell-type-selective trait associations. Conclusions EnSEMBLE shifts transcriptomic interpretation from enumerating possibilities to adjudicating evidence, yielding a compact, traceable set of enhancer-corroborated claims that identify the regulatory programs driving cellular change and prioritize them for experimental validation.

13
AbPACER: parent-aware, affinity-label-blind prioritization of affinity-matured scFv clones from phage-display NGS

Chung, A. J.; Park, B. Y.; Park, E.-B.; Han, J.-H.

2026-08-11 bioinformatics 10.64898/2026.08.05.742956 medRxiv
Top 0.3%
7.9%
Show abstract

BackgroundAffinity-maturation phage-display next-generation sequencing (NGS) yields more paired single-chain variable fragment clones than can be characterized experimentally, creating a fixed-budget prioritization problem. Read counts provide empirical support rather than direct affinity labels. We developed AbPACER (Antibody Parent-Aware Contextual Evidence Ranker), an affinity-label-blind neural ranker combining parent-relative mutation descriptors, frozen antibody-language-model context, and NGS evidence from related clones. AbPACER is campaign-adaptive rather than zero-shot: for each campaign, it is fitted to paired sequences and round-resolved R1-R3 counts before returning a 384-candidate assay list. We evaluated it in two retrospective phage-display campaigns and separately assessed its supervised mean-squared-error adaptation on AlphaSeq, denoted AbPACER-MSE. ResultsFrom frozen top-5% candidate sets containing 16,323 Fas-associated factor 1 (FAF1) and 7,487 vascular endothelial growth factor receptor (VEGFR) clones, each method ranked the complete target-specific set and selected 384 candidates. In FAF1, AbPACER recovered 2.00 {+/-} 0.00 of seven retrospective panel clones, recovering two in every seed, compared with 1/7 by total count, 1.00 {+/-} 0.00 by Ens-Grad CNN, 1.67 {+/-} 1.15 by A2Binder-HL, and 1.33 {+/-} 0.58 by AbAffinity. In VEGFR, AbPACER recovered 2.33 {+/-} 0.58 of three panel clones, the highest observed learned-method mean, whereas total count recovered 3/3. No learned method was uniformly best at broader hypothetical budgets. On the public AlphaSeq common split of 11,670 fixed-test variants, AbPACER-MSE recovered 187.0 {+/-} 2.6 of the true top-384, closely matching AbAffinity (188.0 {+/-} 2.6) and exceeding A2Binder (175.7 {+/-} 6.4) and Ens-Grad CNN (154.0 {+/-} 6.1). AbPACER-MSE updated 1.378 million task-specific parameters, compared with 651.04 million for AbAffinity, and achieved Pearson 0.687 {+/-} 0.003 and Spearman 0.652 {+/-} 0.002. ConclusionsAbPACER provides a campaign-specific, parent-aware framework for fixed-budget prioritization from affinity-label-blind phage-display NGS data. At the 384-candidate endpoint, it showed the highest mean recovery among learned methods in both retrospective campaigns. AbPACER-MSE closely matched AbAffinity in true top-384 recovery while updating substantially fewer task-specific parameters. These results motivate prospective evaluation of sequence-conditioned reranking as a complement to count-based prioritization.

14
SeqDesk: a sequencing-facility management system for standards-compliant and FAIR (meta)data submission

Muench, P. C.; Robertson, G.; McHardy, A. C.

2026-08-11 bioinformatics 10.64898/2026.08.05.743014 medRxiv
Top 0.3%
7.8%
Show abstract

Achieving FAIR compliance requires both standardized metadata and infrastructure for data deposition, yet in practice a large fraction of sequencing studies is still published without the persistent, standards-compliant metadata that reuse depends on. Collecting MIxS-compliant metadata is complex: environment-specific checklists can contain hundreds of fields, and the effort is magnified when metadata is assembled retrospectively at publication time rather than captured throughout the project. We developed SeqDesk, an open-source data management system for sequencing facilities that is designed so that FAIR-compliant public data is produced as the natural output of routine operations. Its current scope is microbial sequencing data, covering metagenomes as well as isolate genomes, for which it supports the corresponding MIxS checklists. SeqDesk gives a sequencing facility a configurable order-and-tracking system for sequencing projects, captures and validates MIxS-compliant metadata aligned with ENA checklists at project initiation, runs bioinformatics analyses through Nextflow pipelines, and brokers submission to the European Nucleotide Archive, all within the institutions own infrastructure. By embedding standards-compliant metadata capture into the sequencing-facility workflow rather than bolting it on at submission, SeqDesk shortens the path from sample to reusable public data. The underlying checklist model is generic, so support can be extended to further data types and metadata standards beyond the microbial domain. SeqDesk is free and open source under the Apache 2.0 licence and available at https://seqdesk.org, with a live demonstration at https://seqdesk.org/#demo.

15
TriTower-m6Am: a triple-tower heterogeneous deep learning architecture integrating semantic, sequential, and structural information for mRNA N6,2'-O-dimethyladenosine site prediction

Xiong, K.; Jia, J.

2026-08-13 bioinformatics 10.64898/2026.08.13.744628 medRxiv
Top 0.3%
7.8%
Show abstract

BackgroundN6,2-O-dimethyladenosine (m6Am) is a cap-proximal mRNA modification deposited by PCIF1 at the first transcribed nucleotide of eukaryotic mRNAs. Knowing where m6Am sites sit across the transcriptome would help explain how cells tune mRNA stability and translation, but current computational predictors typically depend on a single sequence representation and do not jointly model the semantic, sequential, and structural signals carried by an RNA sequence. ResultsWe present TriTower-m6Am, a triple-tower architecture that combines three representations: semantic (RNA-FM with BellPooling), sequential (One-Hot BiLSTM), and structural (RGCN with three typed edges). On an independent test set of 640 sequences, TriTower-m6Am reaches AUC = 0.776, MCC = 0.440, and SN = 0.888, against DTC-m6Ams AUC = 0.765, MCC = 0.411, and SN = 0.800. The 8.8 percentage-point gain in sensitivity means that, for every 100 real m6Am sites, the model recovers roughly 9 additional sites missed by the previous best method. Among the three towers, RGCN alone gives the strongest single signal, and the AUC-weighted ensemble raises sensitivity from the 0.55-0.76 band of the standalone towers to 0.89. Ablating the RGCN edge types shows that backbone connectivity accounts for most of the structural signal. ConclusionsCombining semantic, sequential, and structural views of the same RNA sequence improves m6Am prediction beyond what any single representation achieves. Because each towers contribution to the final prediction is a readable voting weight rather than a hidden parameter, the model is not a black box: a user can read off which tower drove a given prediction and trace it back to the corresponding representation, without running a separate post-hoc explainer. The same design pattern can be transferred to other RNA modification site prediction tasks. Author summaryPredicting where m6Am modifications occur on messenger RNA is important for understanding how cells regulate transcript stability and translation. Existing computational methods typically encode the RNA sequence in a single way, such as k-mer counts or a one-hot code, and treat the model as a black box that emits a prediction without explaining which features drove it. We built TriTower-m6Am to address both limitations. Our model combines three independent encoders--a pretrained RNA language model for semantic patterns, a bidirectional LSTM for local nucleotide order, and a relational graph convolutional network for the structural fold--and fuses their outputs by AUC-weighted voting, so the contribution of each tower to a given prediction is a readable number rather than a hidden parameter. On an independent benchmark the ensemble improves sensitivity by 8.8 percentage points over the prior best method, and ablating the graphs edge types reveals that linear backbone connectivity, rather than long-range base-pairing, carries most of the structural signal. The same triple-tower pattern can be transferred to other RNA modification site prediction tasks.

16
Trust-Aware Sequence-to-Function Modelling in Regulatory Genomics

Onawole, A.; Basiru, S.; Sanni, M. O.; Aiyedun, M.; Sulaimon, R.

2026-08-24 genomics 10.64898/2026.08.20.745945 medRxiv
Top 0.4%
7.6%
Show abstract

Objective: Sequence-to-function models increasingly predict regulatory activity, such as chromatin accessibility, directly from DNA sequence, and are used to interpret non-coding genetic variation. Standard accuracy metrics, computed over a held-out set of genomic regions, do not establish whether an individual prediction remains reliable once the input sequence departs from that set, nor whether a model's attribution-based explanation is biologically grounded rather than coincidental. We develop and evaluate RegTrust-XAI, a trust-aware framework separating these questions using three inference-time signals: ensemble consensus, motif-grounded attribution coherence, and applicability-domain distance. Methods: A five-model convolutional ensemble was trained on 517,790 K562 ATAC-seq windows and evaluated on a held-out chromosome test set (chr8/chr9, n = 42,844). Consensus, coherence, and applicability-domain distance were each tested against prediction error, alongside complementary sequence-novelty analyses and validation against an independent lentiMPRA reporter assay and saturation-mutagenesis MPRA data at the PKLR promoter. Results: The ensemble reached Spearman {rho} = 0.782, with skill of 0.328 over a constant-value null predictor. High-consensus predictions (Scenarios A+B) were consistently enriched for lower error than low-consensus predictions (Scenarios C+D), and attribution coherence further separated error within the high-consensus population (mean absolute error 0.396 versus 0.435, p = 9.6e-10). Applicability-domain distance showed a monotonic error gradient across six distance bands. A 4-mer composition-divergence metric was negatively associated with error and anti-correlated with applicability-domain distance, so composition-based and model-relevant novelty are not equivalent. Attribution transfer to lentiMPRA was assay- and subgroup-dependent, and predicted allele-substitution effects correlated with measured saturation-mutagenesis effects at the PKLR promoter at both 24 h and 48 h ({rho} = 0.227 and 0.235). Motif-specific perturbation further showed that regulatory attributions were strongly context-dependent, with more than 90% of multi-instance motif modules exhibiting superadditive joint effects. Conclusions: Prediction reliability, explanation validity, and sequence novelty are related but distinct properties of a sequence-to-function model. Evaluating each explicitly gives a more complete basis for deciding when to act on a prediction than accuracy alone.

17
A Scalable Framework for Harmonized mtDNA Analysis Across Diverse Biobanks

Schecter, D. R.; Lee, S. S.; Vimal, T.; Lahoti, Y.; Goncalves, V. F.; Retallick-Townsley, K.; Pang, J.; Guvenek, A.; Preuss, M.; Tinker, R. J.; Morava, E.; Kozicz, T.; Hirano, M.; Ganesh, J.; Naini, A.; Liang, J.; Davis, L.

2026-08-25 genetic and genomic medicine 10.64898/2026.08.21.26361041 medRxiv
Top 0.4%
7.1%
Show abstract

Mitochondrial DNA (mtDNA) is increasingly recognized as an important contributor to human disease and population variation, yet most genomic biobanks do not provide standardized mtDNA variant datasets despite abundant mitochondrial sequencing reads in existing whole exome and whole-genome sequencing data. We developed a scalable framework based on the Mitoverse mtDNA Server 2 Fusion workflow to generate harmonized, analysis-ready mtDNA resources across diverse biobank infrastructures. The framework was implemented in the Mount Sinai Million Health Discoveries Program (54,151 participants) using the native Nextflow workflow and adapted for the All of Us Research Program (197,361 participants) using a custom cloud implementation that preserved the same analytical strategy. Across 251,512 participants, the framework generated standardized mtDNA datasets containing 12.9 million variant observations suitable for downstream genomic and electronic health record linked analyses. This framework enables reproducible, population-scale mitochondrial genomics across institutional and national biobanks without requiring additional sequencing or development of new variant calling methods.

18
Benchmarking Imputation Methods for Single-Cell RNA Sequencing Data Using Peripheral Blood Mononuclear Cells from Acute Myocardial Infarction Patients

Ramesh, P.; Fyta, M.

2026-08-27 bioinformatics 10.64898/2026.08.23.746230 medRxiv
Top 0.4%
7.0%
Show abstract

Acute myocardial infarction (AMI) remains one of the leading causes of mortality worldwide, and the following post-effects, such as post-AMI inflammation and tissue repair, involve peripheral blood mononuclear cells playing a critical role. The influence of imputation methods in biological data is assessed with respect to high-resolution single-cell RNA sequencing (scRNAseq) data relevant to these cells. Still scRNAseq data often encounter a lot of dropout events, leading to sparse and noisy datasets, hampering downstream results. To assess the influence of the missingness in the data, we artificially impose different levels of dropout in available scRNAseq data by leveraging various imputation techniques. Specifically, we introduce artificial missingness at 10%, 20%, and 30% levels under a missing completely at random (MCAR) framework, repeated across 10 independent runs. We benchmarked six imputation strategies - MAGIC, IterativeImputer, KNNImputer, Mean Imputation, SoftImpute, and a Generative adversarial network (GAN) - based approaches using multiple evaluation metrics: marker gene preservation, clustering consistency (Adjusted Rand Index - ARI), gene-wise correlation with ground truth, and structural separation (silhouette scores). The results clearly underline that no single imputation method dominated across all metrics. Overall, Mean and KNN imputers showed limited recovery across all benchmarks. GAN excelled in global transcriptional recovery and SoftImpute in preserving biologically meaningful cell-type signals. Our results highlight the importance of selecting the imputation methods as part of the pre-processing step towards the downstream biological questions related to transcriptome recovery, detection of marker genes, or maintaining cell-type-specific resolution.

19
Correction of the cytosine deamination artifacts in FFPE-based sequencing experiments

Płonka, W.; Kostka, D.; Lalik, A.; Kurpas, M.; Dinh, K. N.; Sitkiewicz, M.; Kimmel, M.; Rzyman, W.; Jaksik, R.

2026-08-19 bioinformatics 10.64898/2026.08.11.744151 medRxiv
Top 0.4%
6.8%
Show abstract

Formalin-fixed, paraffin-embedded (FFPE) tissues remain an essential resource for molecular studies, yet formalin-induced cytosine deamination introduces characteristic C>T/G>A artifacts that compromise the accuracy of next-generation sequencing (NGS) analyses. Numerous computational methods and enzymatic DNA repair strategies have been proposed to reduce these artifacts, but no systematic comparison across tools and experimental conditions exists. Here, we evaluate the performance of seven computational approaches (SOBDetector, Ideafix, MicroSEC, FFPolish, DeepOmics FFPE/FFPE-PLUS, FFPErase) together with the NEBNext(R) FFPE DNA Repair Mix v2, a multi-enzyme repair system applied during DNA preparation. Using three independent datasets, one based on whole genome sequencing (CGCI-BL) and two on whole exome sequencing (TCGA-PC and SUT-LUAD, the latter containing enzymatically repaired samples), and matched fresh-frozen samples as the gold standard, we assess precision, sensitivity, and artifact reduction efficiency across all methods. We further examine the potential synergy between enzymatic repair and post-sequencing computational filtering. Our results provide practical guidelines for FFPE artifact correction and demonstrate that enzymatic treatment provides the best results, while among the computational methods, FFPErase offers the most robust reduction of cytosine deamination artifacts while maximizing the retention of true somatic variants. KEY MESSAGESO_LIFormalin fixation in FFPE samples introduces artifacts that can significantly affect the accuracy of NGS analyses. C_LIO_LIAmong the evaluated approaches, enzymatic repair using NEBNext(R) FFPE DNA Repair Mix v2 achieves the most effective reduction of sequencing artifacts. C_LIO_LIComputational methods vary in performance, with FFPErase showing the most robust balance between artifact removal and retention of true somatic variants. C_LIO_LICombining enzymatic repair with computational filtering did not lead to consistent improvements in performance across datasets. C_LI

20
A comprehensive benchmark of transcriptome-wide fusion detection using long-read RNA sequencing

Dorney, R.; Wu, S.; Hung, J. Y.-H.; Hebbard, L.; Schmitz, U.

2026-08-12 bioinformatics 10.64898/2026.08.07.743439 medRxiv
Top 0.4%
6.7%
Show abstract

Fusion transcripts contribute to cancer, inherited diseases, developmental disorders, and evolution. Long-read RNA sequencing enables direct sequencing of full-length transcripts, creating new opportunities to detect complex fusion architectures, including previously inaccessible multi-segmented fusion transcripts. However, accurate transcriptome-wide fusion detection remains challenging because existing methods struggle to distinguish genuine fusion events from technical artefacts. Here, we present a comprehensive benchmark of transcriptome-wide fusion detection using simulated datasets and transcriptomes from three cancer cell lines across Oxford Nanopore Technologies (ONT) cDNA, PCR-cDNA, and direct RNA sequencing, Pacific Biosciences (PacBio) Kinnex sequencing, Illumina short-read RNA sequencing, six long-read fusion callers, and multiple analysis strategies. False-positive fusion calls remained the dominant limitation across sequencing platforms and algorithms. Increasing sequencing depth improved recall but also amplified spurious fusion calls, whereas higher read-support thresholds improved precision at the expense of sensitivity. ONT PCR-cDNA sequencing combined with CTAT-LR-Fusion achieved the best overall balance between precision and recall, whereas JAFFAL was the only caller to reliably identify simulated tri-gene fusions. Consensus calling reduced false positives but markedly reduced sensitivity, with only one of 400 simulated fusions detected by all six callers. Breakpoint localisation emerged as a major limitation across all methods. Long-read sequencing consistently recovered more validated fusion transcripts than short-read sequencing, enabled detection of complex tri-gene fusions, and produced more biologically plausible fusion landscapes with fewer promiscuous gene partners. Collectively, our results establish the first comprehensive benchmarking framework for transcriptome-wide fusion detection, using long-read RNA sequencing, and provide practical guidance for selecting sequencing workflows and computational strategies, while identifying key priorities for future algorithm development.